Verbal Morphosyntactic Disambiguation through Topological Field Recognition in German-Language Law Texts

نویسندگان

  • Kyoko Sugisaki
  • Stefan Höfler
چکیده

The morphosyntactic disambiguation of verbs is a crucial pre-processing step for the syntactic analysis of morphologically rich languages like German and domains with complex clause structures like law texts. This paper explores how much linguistically motivated rules can contribute to the task. It introduces an incremental system of verbal morphosyntactic disambiguation that exploits the concept of topological fields. The system presented is capable of reducing the rate of POS-tagging mistakes from 10.2% to 1.6%. The evaluation shows that this reduction is mostly gained through checking the compatibility of morphosyntactic features within the long-distance syntactic relationships of discontinuous verbal elements. Furthermore, the present study shows that in law texts, the average distance between the left and right bracket of clauses is relatively large (9.5 tokens), and that in this domain, a wide context window is therefore necessary for the morphosyntactic disambiguation of verbs. DOI: https://doi.org/10.1007/978-3-642-40486-3_8 Posted at the Zurich Open Repository and Archive, University of Zurich ZORA URL: https://doi.org/10.5167/uzh-79353 Accepted Version Originally published at: Sugisaki, Kyoko; Höfler, Stefan (2013). Verbal morphosyntactic disambiguation through topological field recognition in German-language law texts. In: Mahlow, Cerstin; Piotrowski, Michael. Systems and Frameworks for Computational Morphology. Berlin Heidelberg: Springer, 136-147. DOI: https://doi.org/10.1007/978-3-642-40486-3_8 Verbal Morphosyntactic Disambiguation through Topological Field Recognition in German-Language Law Texts Kyoko Sugisaki and Stefan Höfler⋆ University of Zurich, Institute of Computational Linguistics, Binzmühlestrasse 14, 8050 Zürich, Switzerland {sugisaki,hoefler}@cl.uzh.ch http://www.cl.uzh.ch Abstract. The morphosyntactic disambiguation of verbs is a crucial pre-processing step for the syntactic analysis of morphologically rich languages like German and domains with complex clause structures like law texts. This paper explores how much linguistically motivated rules can contribute to the task. It introduces an incremental system of verbal morphosyntactic disambiguation that exploits the concept of topological fields. The system presented is capable of reducing the rate of POS-tagging mistakes from 10.2% to 1.6%. The evaluation shows that this reduction is mostly gained through checking the compatibility of morphosyntactic features within the long-distance syntactic relationships of discontinuous verbal elements. Furthermore, the present study shows that in law texts, the average distance between the left and right bracket of clauses is relatively large (9.5 tokens), and that in this domain, a wide context window is therefore necessary for the morphosyntactic disambiguation of verbs. The morphosyntactic disambiguation of verbs is a crucial pre-processing step for the syntactic analysis of morphologically rich languages like German and domains with complex clause structures like law texts. This paper explores how much linguistically motivated rules can contribute to the task. It introduces an incremental system of verbal morphosyntactic disambiguation that exploits the concept of topological fields. The system presented is capable of reducing the rate of POS-tagging mistakes from 10.2% to 1.6%. The evaluation shows that this reduction is mostly gained through checking the compatibility of morphosyntactic features within the long-distance syntactic relationships of discontinuous verbal elements. Furthermore, the present study shows that in law texts, the average distance between the left and right bracket of clauses is relatively large (9.5 tokens), and that in this domain, a wide context window is therefore necessary for the morphosyntactic disambiguation of verbs.

برای دانلود متن کامل این مقاله و بیش از 32 میلیون مقاله دیگر ابتدا ثبت نام کنید

ثبت نام

اگر عضو سایت هستید لطفا وارد حساب کاربری خود شوید

منابع مشابه

Incremental Morphosyntactic Disambiguation of Nouns in German-Language Law Texts∗

Morphosyntactic disambiguation is a crucial pre-processing step for the recognition of grammatical functions in morphologically rich languages like German and heavily nominalized domains like law texts. This paper explores how far linguistically motivated hard rules can contribute to morphosyntactic disambiguation. It introduces an incremental system that is capable of reducing the rate of morp...

متن کامل

Implementation of Croatian NERC System

In this paper a system for Named Entity Recognition and Classification in Croatian language is described. The system is composed of the module for sentence segmentation, inflectional lexicon of common words, inflectional lexicon of names and regular local grammars for automatic recognition of numerical and temporal expressions. After the first step (sentence segmentation), the system attaches t...

متن کامل

The Philosophy and Functions of Verbal Violence in Harold Pinter’s Mountain Language: A CDA Approach

The present study considers the issue of verbal violence in the language of drama. In the evaluation of verbal violence, Jeanette Malkin (2004) proposes six maxims, through which language may be considered as an arrogant element. The characters in dramatic texts (as in other literary texts) are created, developed, evolved and - in some cases - destroyed by language. In a considerable numb...

متن کامل

Combining Morphosyntactic Enriched Representation with n-best Reranking in Statistical Translation

The purpose of this work is to explore the integration of morphosyntactic information into the translation model itself, by enriching words with their morphosyntactic categories. We investigate word disambiguation using morphosyntactic categories, n-best hypotheses reranking, and the combination of both methods with word or morphosyntactic n-gram language model reranking. Experiments are carrie...

متن کامل

The Impact of Semantic and Morphosyntactic Ambiguity on Automatic Humour Recognition

Humour is one of the most amazing characteristics that defines us as human beings and social entities. Its study supposes a deep insight into several areas such as linguistics, psychology or philosophy. From the Natural Language Processing (NLP) perspective, recent researches have shown that humour can be automatically generated and recognized with some success. In this work we present a study ...

متن کامل

ذخیره در منابع من


  با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید

برای دانلود متن کامل این مقاله و بیش از 32 میلیون مقاله دیگر ابتدا ثبت نام کنید

ثبت نام

اگر عضو سایت هستید لطفا وارد حساب کاربری خود شوید

عنوان ژورنال:

دوره   شماره 

صفحات  -

تاریخ انتشار 2013